Back

Communications Chemistry

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match Communications Chemistry's content profile, based on 48 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.

1
Leveraging AI and structural proteomics for rational design of a KAT6A degrader

Arad, G.; Simchi, N.; Brodsky, S.; Shtrikman, A.; Kedem, Y.; Alchanati, I.; Otonin, G.; Shenoy, A.; Kovalerchik, D.; Ran Shchory, M.; Ben Shoshan-Galeczki, Y.; Cohen, N.; Lange, K.; Seger, E.; Pevzner, K.

2026-05-28 bioinformatics 10.64898/2026.05.25.727609 medRxiv
Top 0.1%
18.3%
Show abstract

While targeted protein degraders such as PROTACs are a clinically proven therapeutic strategy, the discovery of novel degraders remains hampered by trial-and-error process. To address this challenge, we developed the AIMS platform, which combines structural proteomics with AI models for rational PROTAC design. AIMS is an end-to-end toolkit for PROTAC optimization, encompassing structure solving using proteomics and AI, prediction of ADME and degradation properties, and prospective ranking of compound design ideas. Altogether, this integrated platform successfully enabled the multi-parameter optimization of a potent and bioavailable in vivo validated KAT6A degrader, establishing a versatile framework for PROTAC development across various targets. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=71 SRC="FIGDIR/small/727609v1_ufig1.gif" ALT="Figure 1"> View larger version (17K): org.highwire.dtl.DTLVardef@13596d0org.highwire.dtl.DTLVardef@140500eorg.highwire.dtl.DTLVardef@147e585org.highwire.dtl.DTLVardef@12dbdfe_HPS_FORMAT_FIGEXP M_FIG C_FIG

2
Room-temperature fragment screening of soluble epoxide hydrolase by serial crystallography

Dunge, A.; Wehlander, G.; Branden, G.; Kack, H.

2026-06-02 biochemistry 10.64898/2026.06.01.729266 medRxiv
Top 0.1%
18.2%
Show abstract

Room temperature serial crystallography offers advantages over conventional cryo-crystallography, such as simplified crystal handling and the possibility to avoid potential artefacts associated with cryo-trapping. However, to be considered as an alternative for drug discovery, where compound availability may be limited and speed of structure delivery is a key factor, it suffers from several limitations. To address these challenges, we have optimized a serial crystallography workflow for ligand soaking, data collection and data processing, significantly reducing time and reagent consumption to make it a viable option for drug discovery applications, herein exemplified by crystallographic fragment screening. Our approach incorporates the use of dried-in fragment cocktails on fixed target supports, compatible with 96-well plates for crystal soaking, and an efficient data processing pipeline tailored for serial crystallography. To validate our workflow, we conducted an in-crystal fragment screen at room temperature on the protein soluble epoxide hydrolase. The screen comprised 384 compounds and resulted in identification of 40 fragment binders corresponding to a hit rate of 10.4 %. The resulting room-temperature structures are of high quality and reveal opportunities for specific interaction within the highly hydrophobic active site of soluble epoxide hydrolase. Finally, we discuss potential avenues for further workflow optimization, highlighting the future potential of this approach for drug discovery. SynopsisWe have developed a workflow that allowed us to efficiently conduct a fragment screen at room temperature using serial crystallography, of interest for future drug discovery campaigns.

3
Chemoinformatics-guided discovery of food-grade anionic stabilizers for phycocyanin under acidic conditions

law, l.; Chuang, K.; Luo, L.

2026-05-27 biochemistry 10.64898/2026.05.25.727568 medRxiv
Top 0.1%
12.5%
Show abstract

Phycocyanin (PC) is the principal natural blue pigment used in functional beverages, but it rapidly loses color and aggregates under acidic conditions (pH {approx} 3). Experimental screening of stabilizers is costly and combinatorially intractable. Here we develop a chemoinformatics framework -- descriptor-based QSPR, a chemistry-prior heuristic, and virtual screening -- that learns from three rounds of commissioned screening (48 compounds, 6% hit rate) to predict stabilizer efficacy directly from molecular structure. In this genuinely small-data regime (3 positives), a LightGBM classifier built from 10 RDKit descriptors and 11 domain-expert charge/polymer features attained a leave-one-out AUC of 0.73, only marginally above a single-feature charge-density baseline (AUC 0.67); LOO sensitivity was 1/3 at threshold 0.5. A complementary chemistry-prior heuristic encoding anion-type priors substantially outperformed both, reaching AUC 0.95, indicating that explicit chemical knowledge captures information that descriptor-based ML cannot readily recover at this dataset size. SHAP analysis of the QSPR identified effective negative-charge density per unit, log molecular weight, polyphosphate identity, and functional-group density as the dominant features (jointly {approx}97% of mean |SHAP|), recovering the electrostatic-complexation mechanism without it being supplied as a prior. Virtual screening of 30 generally recognized as safe (GRAS) food additives nominated the pyrophosphate family -- led by sodium pyrophosphate decahydrate and sodium hexametaphosphate (SHMP), both at P {approx} 0.99 -- and the heuristic additionally flagged sodium phytate (IP{square}), which the descriptor model under-ranked at P = 0.045. Experimental validation at pH 3 and 46 {degrees}C for 7 days confirmed SHMP 2:1 (78.1 {+/-} 11.3% color retention), TSPP 2:1 (54.1 {+/-} 10.6%) and IP{square} 1:1 (52.7 {+/-} 9.0%), while a ternary IP{square} + STPP combination reached 83.8 {+/-} 11.9%, surpassing all single-component formulations. {zeta}-Potential measurements indicated a predominantly electrostatic origin for the protection (Pearson r = -0.82 between {zeta} and CR{square} {square} {square}; n = 24; p = 1 x 10{square} {square}). The framework, dataset and code are released to accelerate stabilizer discovery for other acid-sensitive food colorants and to provide a candid small-data benchmark.

4
Learning molecular determinants of selective small-molecule partitioning across biomolecular condensates

Khambhawala, A.; Rekhi, S.; Chen, Q.; Mohanty, P.; Tabor, D. P.; Mittal, J.

2026-05-27 biochemistry 10.64898/2026.05.23.727301 medRxiv
Top 0.1%
11.8%
Show abstract

The functional role of biomolecular condensates is shaped by the composition of constituent proteins, nucleic acids, ions, and small molecules. Selective partitioning of small molecules into condensates has therefore emerged as a potential route to condensate-specific chemical probes and therapeutics. Although partitioning is influenced by differences in solvation environments between coexisting dense and dilute phases, a molecular framework connecting small-molecule structure to condensate-specific enrichment remains lacking. Here, we use existing experimental partitioning data for a library of FDA-approved drugs and metabolites across four biomolecular condensates to develop an interpretable graph-based model of small-molecule partitioning. By combining multitask pretraining, condensate-specific fine-tuning, evidential uncertainty quantification, and atom-level attribution analysis, our model predicts continuous partition coefficients with improved accuracy over descriptor-based approaches. Atom-level attributions reveal that condensate partitioning is not governed by a universal chemical rule: the same molecular scaffold can be read differently by distinct condensate environments, with local atomic context and connectivity determining whether specific atoms promote or suppress enrichment. We further apply the trained model to ~1.7 million drug-like molecules from ChEMBL, identifying a chemically diverse space of predicted condensate-selective partitioners and mapping regions where predictions are confident versus where new measurements would be most informative. Together, this work establishes condensate partitioning as a chemically learnable property shaped by the interplay between small-molecule structure and condensate-specific microenvironments, providing an interpretable and uncertainty-aware framework for defining molecular determinants of partitioning and guiding the discovery of condensate-selective small molecules.

5
A tunable aqueous architecture modulates functionaloutput in biomolecular condensates

Sasazawa, M.; Chen, M.; Zeng, R.; Denis, U.; Bais, S.; Hoffstadt, J.; von Hofe, J.; Hoffmann, N.; Volkova, Y.; Saurabh, S.

2026-05-15 biochemistry 10.64898/2026.05.13.724666 medRxiv
Top 0.1%
11.8%
Show abstract

Biomolecular condensates organize cellular biochemistry, yet the principles governing their internal solvent architectures remain poorly understood. Most current models focus on macromolecular scaffolds while treating the solvent as a passive, spatially uniform background. Here, we introduce Condensate Spatial Topography via Emission Lifetimes (ConSTEL) to map the continuous solvent polarity landscape inside biomolecular condensates. Using PopZ as a model system, we show that the condensate interior contains a persistent, tunable mosaic of aqueous environments whose apparent polarity, reported by Nile Red fluorescence lifetimes, is organized by thermodynamic state and chemical cues. This microphase-separated solvent architecture defines distinct mesoscale rheological regimes, with intermediate aqueous niches supporting fast, confined tracer motion and highly polar or non-polar extremes forming a slower, viscoelastic mesh. We further demonstrate that drug-like small molecules partition non-uniformly across this landscape according to their physicochemical properties, and that exceeding local solubility limits drives "reciprocal sculpting", in which mismatched guests remodel the host solvent architecture. Together, these results highlight internal solvent organization as an active, tunable determinant of condensate material properties, molecular transport, and partitioning, and suggest that predictive models of condensate function and pharmacology would benefit from incorporating the spatial arrangement of solvent environments alongside bulk composition.

6
Effects of PTMs on Tau Protein Aggregation: Insights from HCG and Atomistic MD Simulations

Louet, A. A. B.; Stuke, J.; Pietrek, L.; Vendruscolo, M.; Hummer, G.

2026-05-26 biochemistry 10.64898/2026.05.22.727278 medRxiv
Top 0.1%
11.6%
Show abstract

Post-translational modifications (PTMs) of the tau protein are increasingly recognized as pivotal regulators in the onset and progression of tauopathies, such as Alzheimers disease (AD). To systematically evaluate the structural and functional consequences of specific PTMs, we generated and analyzed seven distinctly modified variants of the tau-K32 construct. These included phosphorylation at Ser202/Thr205, phosphorylation at Ser258/Ser262/Ser356, full phosphorylation at all reported Ser/Thr sites, acetylation at Lys274/Lys281, acetylation at Lys280, full acetylation at all sites, and an unmodified control. Selection of PTM sites was guided by prior experimental literature. By incorporating fully modified tau models, we assessed the global impact of widespread modifications on structural properties and aggregation behavior. Our findings establish a comparative framework for understanding how discrete and cumulative PTMs modulate tau aggregation and provide mechanistic insight into PTM-induced tau dysfunction relevant to neurodegenerative diseases.

7
Instance-Wise Contrastive Graph Neural Network Enables the Discovery of Novel Aedes aegypti Larvicidal Compounds

da Costa, K. S. L.; Caldeira, G. H. G.; Costa, V. A. F.; Silva, A. S.; Pereira, C. d. S.; Batista, B. C.; Manchein, L. B.; Martin, H.-J.; Rafique, J.; Braga, R. d. C.; Muratov, E.; Saba, S.; de Oliveira, G. A. R.; Luz, C.; Rodrigues, J.; Neves, B. J.

2026-05-31 bioinformatics 10.64898/2026.05.28.726277 medRxiv
Top 0.1%
11.1%
Show abstract

Aedes aegypti remains a major arboviral vector, making larval control a critical strategy to reduce mosquito populations. However, resistance to commercial larvicides has reduced the long-term effectiveness of current interventions, reinforcing the need for new compounds with improved potency and selectivity. Here, we present an instance-wise contrastive graph neural network (GNN) framework to accelerate the discovery of novel larvicidal compounds. The model was trained on a curated dataset of 556 organic compounds organized into LC50-derived multitask classification thresholds and integrated Transformer-inspired graph learning with whole-molecule and fragment-level contrastive regularization. This model achieved strong held-out performance, with global AUC = 0.95 {+/-} 0.01, PR-AUC = 0.93 {+/-} 0.01, and MCC = 0.77 {+/-} 0.03, outperforming conventional machine learning and graph-based baselines. Predictive uncertainty analysis and counterfactual maps further supported the interpretation of threshold-sensitive predictions and substructural contribution patterns. The model was applied to screen 1.3 million compounds, resulting in 10 candidates for experimental validation. Three compounds showed measurable larvicidal activity against A. aegypti larvae. Among them, LC-79 emerged as the most promising hit, with 2-day and 5-day LC50 values of 0.24 {micro}g/mL (0.66 {micro}M) and 0.05 {micro}g/mL (0.13 {micro}M), respectively, an IE50 of 0.06 {micro}g/mL (0.16 {micro}M), and rapid larval mortality (LT50 = 1.10 days at 1 {micro}g/mL). LC-79 also showed no measurable acute toxicity to Daphnia magna at the highest tested concentration [EC50-48h >43 {micro}g/mL (>119 {micro}M)], resulting in selectivity indices >180 and >860 relative to its 2-day and 5-day LC50 values. Overall, this study demonstrates that contrastive graph learning can move beyond retrospective larvicide modeling to experimentally validated hit discovery, identifying LC-79 as a potent and preliminarily selective acylthiourea larvicide candidate for further mechanism-of-action, resistance, and semi-field evaluation.

8
Function-guided design of active enzymes

Hu, M.; Wu, L.; Yang, Y.; Li, F.; Zhu, L.

2026-06-29 bioinformatics 10.64898/2026.06.27.735025 medRxiv
Top 0.1%
11.0%
Show abstract

Designing enzymes from functional descriptions remains challenging because catalytic activity is governed by sequence-structure-function relationships. Here we present EnzymeArt, a function-conditioned enzyme-design framework centred on a generative sequence model. EnzymeArt couples function-conditioned sequence generation with structure-guided refinement, annotation checks and substrate-aware computational prioritization to select candidates for synthesis and biochemical testing. Across alcohol dehydrogenase (ADH), malate dehydrogenase (MDH) and triacylglycerol lipase design campaigns, 57 of 60 synthesized designs showed crude-lysate activity above matched background controls. Purified representatives further showed quantitative steady-state catalytic activity. The best designed ADH reached kcat = 223.7/s and exceeded a wild-type reference under matched conditions, an MDH reached kcat = 267.57/s despite having only 33% sequence identity to its closest BLASTP hit, and a designed lipase hydrolysed both short- and long-chain triglycerides with apparent activity modestly above that of a commercial lipase reference. Together, these results establish a route for converting functional descriptions into experimentally validated enzyme designs with quantitative steady-state kinetic activity.

9
Phase composition-specific behaviour of functional RNAs in liquid-liquid phase-separated microenvironment

Chakraborty, A.; Khan, F.; Sharma, S.; Ameta, S.

2026-05-21 evolutionary biology 10.64898/2026.05.19.726130 medRxiv
Top 0.1%
10.9%
Show abstract

The internal dynamics of liquid-liquid phase-separated systems are governed primarily by polymer packing, excluded-volume effect, and interactions between polymers and encapsulated macro-molecules. Although one immediate effect of such a constrained microenvironment is diffusion limitation, it remains unclear whether encapsulated macromolecules can also exhibit phase composition-specific functional behaviour that is not observable in a well-mixed aqueous environment. In this regard, different phases in a phase-separated environment can be accessed via a phase diagram that demarcates the region between two-phase (droplets) and one-phase (polymer-rich, no droplets) regimes. While the two-phase region is heterogeneous, most previous work on encapsulating functional macromolecules in phase-separated droplets uses a single point from the phase diagram. This leaves a clear gap in understanding on how the function scales across this landscape of droplets and identifying regions advantageous for the encapsulated macromolecule and its function. Here, using the Spinach light-up RNA aptamer, we show that RNA function does not scale uniformly across the phase diagram. We show that RNA can exhibit phase composition-specific functional behaviour due to constraints imposed by the internal microenvironment of phase-separated droplets. Furthermore, using variants of the Spinach aptamer, we show that fluorescence activity differences among the variants vary differently with phase-separation regimes across the phase map, suggesting that some regions of the phase diagram can confer a selective advantage. Our results highlight the potential of liquid-liquid phase-separated internal microenvironments in guiding the differentiation of functional RNA variants, which could serve as a physical selection pressure in pre-cellular evolution.

10
Divergent Inclusion Body Structures and Stabilities Emerge from Native Monomer Properties

Siebeneichler, B.; Liu, X.; Rodriguez Cruz, P. E.; Naser, D.; DelMistro, G.; Steckner, J.; Schaefer, A.; Tran, N.; Holyoak, T.; Meiering, E. M.

2026-06-02 biochemistry 10.64898/2026.06.01.729379 medRxiv
Top 0.1%
10.8%
Show abstract

Protein aggregation is of broad importance in biotechnology and disease, yet the structural heterogeneity of cellular aggregates has confounded high-resolution structural analysis. Inclusion bodies (IBs) formed in Escherichia coli are an attractive, controllable system for unravelling the complexities of protein aggregation in a cellular context. Here, a multimodal analysis integrating residue-resolved quenched amide hydrogen-deuterium exchange (qHDX), proteolysis, FTIR, Congo red binding, and chemical denaturation is applied to IBs formed by proteins encompassing stable {beta}- and -globular folds, a partially structured protein fragment, and intrinsically disordered low complexity domains (LCDs). Remarkable conformational diversity is observed: IBs formed by well-folded proteins are extensively structured and include substantial local native-like features, whereas proteins with decreased access to stable native conformations form more heterogeneous and dynamic aggregates increasingly shaped by intrinsic sequence features. Strikingly, qHDX protection of TDP-43 LCD IBs strongly aligns with the core of cryo-EM structures of ex vivo pathological fibrils; however, peripheral regions that appear fully hydrogen-bonded in the cryo-EM structures exhibit little protection. The results reveal that individual protein IBs contain distinct mixtures of native-like, disordered, and amyloid-like conformers, informing the prediction and control of cellular aggregate structure and stability.

11
From APOE Genetics to AI-Designed Drug Candidates: An Integrated Pipeline for Oral, Brain-Penetrant ACAT1 Inhibitors in Alzheimer's Disease

Agarwal, S.; Popert, R.; Agapow, P.; Ruff, C.; Gupta, S.

2026-05-26 neuroscience 10.64898/2026.05.21.727002 medRxiv
Top 0.1%
8.9%
Show abstract

For more than 55 million people living with dementia worldwide, no oral disease-modifying treatment is currently available. Alzheimers disease (AD) remains one of the most urgent unmet needs in neurology, with recent genetic evidence estimating that 72-93% of AD burden is attributable to common APOE allelic variation. Mechanistically, the APOE {varepsilon}4 isoform impairs cholesterol transport in the brain, promoting cholesteryl ester accumulation in microglia; these lipid-laden cells lose phagocytic capacity for amyloid-{beta} clearance and autophagy-mediated degradation of phosphorylated tau, linking a single upstream metabolic disruption to both hallmark pathologies of AD. ACAT1/SOAT1, the brains predominant cholesterol-esterifying enzyme, is an attractive therapeutic target: its inhibition reduces cholesteryl ester formation, restoring microglial A{beta} clearance, reducing A{beta} production, and promoting tau degradation via autophagy, supporting a multimodal mechanism not addressed by currently approved therapies. However, prior ACAT inhibitor programs were limited by isoform non-selectivity, excessive lipophilicity, and lack of CNS optimization. Addressing these constraints, we present the CuraGenAI Drug Discovery Platform for the design of oral, brain-penetrant ACAT1-targeted small-molecule candidates. The pipeline integrates scaffold-constrained generative chemistry with multi-objective ADMET optimization across more than 30 criteria. From approximately 6 million generated molecules, multi-stage filtering yielded approximately 7,300 CNS-optimized candidates with predicted oral brain-penetrant profiles distinct from prior ACAT clinical compounds. Three nominated lead candidates showed predicted BBB probabilities >0.93, favorable CNS drug-like physicochemical profiles, and stable ACAT1 binding poses, and are prioritized for synthesis and in vitro validation.

12
Minimal Data, Maximal Insight (MDMI): A Structure-guided Pipeline for Discovering Functional Alternatives in Peptide-Protein Interfaces

Bayat, P.; Perkins, S. J.; Clancy, S.; Patel, S. S.; Yin, R. F.; Bozovicar, K.; Singh, S.; Shrestha, S.; Moustafa, Z.; Zayani, R.; IWE, I.; Bayat, S.; Kelly, P.; Vigar, J. R. J.; White, V. Y.; Xie, M.; Simchi, M.; Palter, S.; Nguyen, J.; Zeisler, I. Y.; Wu, B.; Pardee, K.

2026-07-14 synthetic biology 10.64898/2026.07.13.737974 medRxiv
Top 0.1%
8.5%
Show abstract

Discovering functional peptides across vast sequence space remains a formidable challenge, particularly when experimental training data is scarce. We present Minimal Data Maximal Insight (MDMI), a two-stage structure-guided computational pipeline that designs functional peptide variants using only a small, annotated dataset. Rather than relying on sequence information alone, MDMI integrates three-dimensional structural features derived from predicted peptide-protein complexes into a machine learning model that captures interface geometry and binding energetics. This structure-aware predictor, paired with a genetic algorithm for sequence exploration, reduced false positives from 70% to close to zero in an all-negative benchmark panel compared with a sequence-only model in computational benchmarking, and produced approximately four-fold more high-confidence in silico binders than state-of-the-art peptide/protein design baselines. Using the split-GFP system as a testbed, where fluorescence provides a direct functional readout of peptide-protein complementation, MDMI identified peptides with up to 38% sequence divergence from wild-type in Stage 1 while retaining measurable activity. In Stage 2, motif-guided recombination of successful Stage 1 variants produced highly divergent yet functional peptides bearing over 50% sequence difference from wild-type, revealing two distinct functional clusters in sequence space. As further validation, a top-performing candidate expressed as a full-length GFP fusion retained a GFP-like emission profile, supporting formation of a fluorescent GFP-like scaffold. These results demonstrate that structure-informed pipelines can uncover remote functional sequence space from minimal data, with broad implications for peptide and therapeutic analog discovery.

13
From biofilms to birth: Quantitative murburn rationale for hydrated polymer-centred biological transduction, coherence, and evolution of complex life

Manoj, K. M.; Jaeken, L.; Tamagawa, H.; Burra, V. L. S. P.

2026-07-13 biochemistry 10.64898/2026.07.10.737745 medRxiv
Top 0.1%
8.1%
Show abstract

Hydrated extracellular polymeric phases (such as mucus, biofilms, and extracellular matrices) have traditionally been viewed as passive barriers. We complement and extend this view by analysing these systems through the murburn framework and liquid-liquid phase separation (LLPS) biophysics. Using quantitative modeling, we first demonstrate how frothy mucus in amphibian egg-masses enhances oxygen delivery while buffering diffusible reactive species (DRS), leading to improved developmental synchrony. We then model the human cervical mucus system, showing its cycle-dependent transitions between coherent barriers (pregnancy), active transduction media (ovulation), and controlled inflammatory remodeling (labor). Finally, we present thiolated polyglycerol sulfate (dPGS-SH) as a synthetic validation (another groups recently published work): this rationally designed mucolytic agent recapitulates native mucuss DRS-modulating properties and shows superior efficacy for addressing cystic fibrosis pathology. With such pan-systemic perspectives, we argue that phase-separated hydrated polymeric matrices represent one of evolutions most conserved solutions for regulating stochastic murburn chemistry, enabling organisms to exploit oxygen while preserving biological coherence. From biofilms to birth, this framework unifies the physicochemical basis of lifes most fundamental processes.

14
Multi-Modal Deep Learning Integrates Spatial Topologies and Sequential Motifs to Identify Class I HDAC Inhibitors as Pan-Cancer Therapeutics

Tong, S.; Zhang, W.; Ji, S.

2026-04-25 bioinformatics 10.64898/2026.04.22.720196 medRxiv
Top 0.1%
8.1%
Show abstract

The molecular characterization of human solid tumors has introduced immense genomic complexity and intra-tumoral diversification. Converting these detailed, multi-omic profiles right into workable, broad-spectrum therapeutics continues to be an formidable bottleneck in precision oncology. Traditional computational drug repurposing strategies largely rely on single-modality chemical descriptors, which frequently fail to capture the systemic transcriptomic interactions within the highly dynamic tumor microenvironment. Here, this study presents a robust multi-modal deep learning framework that synergistically integrates two-dimensional (2D) molecular graphs via Graph Neural Networks (GNNs) and chemical functional group patterns via self-attention Transformers. By mapping this dual-stream chemical feature space to the perturbational transcriptomic signatures (LINCS L1000) of 22 distinct cancer types from The Cancer Genome Atlas (TCGA), a vast library of over 28,000 small-molecule compounds was computationally screened. The developed multi-modal architecture achieved state-of-the-art predictive accuracy, significantly outperforming traditional single-modality baseline models. Strikingly, the comprehensive pan-cancer transcriptomic reversal landscape identified a persistent convergence of non-oncology drugs exhibiting potent broad-spectrum anti-tumor potential. Specifically, Class I Histone Deacetylase (HDAC) inhibitorsmost notably TC-H-106, RG2833, and Tianeptinaline, agents originally developed to penetrate the blood-brain barrier for neurodegenerative and psychiatric disordersemerged as top therapeutic candidates across lung adenocarcinoma (LUAD), bladder urothelial carcinoma (BLCA), and rectum adenocarcinoma (READ). Subsequent high-dimensional network pharmacology and functional enrichment analyses confirmed that these agents robustly suppress essential oncogenic pathways, specifically collapsing the G1/S phase transition and DNA damage repair machineries. Furthermore, structural validation via molecular docking and force-field thermodynamics confirmed the highly stable physical binding affinity (Vina score: -7.0 kcal/mol, MMFF94 Energy: 64.76 kcal/mol) of TC-H-106 to the HDAC1 catalytic pocket. Kaplan-Meier survival analysis based on TCGA gene expression stratification underscored the significant prognostic benefit of targeting this epigenetic axis. Collectively, these findings introduce a powerful multi-modal AI framework for systems-level drug repurposing and highlight brain-penetrant Class I HDAC inhibitors as highly promising candidates for pan-cancer epigenetic therapy.

15
Optical Screening Identifies Chemical Modulators of Intracellular α-synuclein Aggregation

Rothschild, L.; Giem, C.; Bajaj, A.; Luo, J. W.; Carey, K. L.; Deguine, J.; Xavier, R. J.

2026-07-10 cell biology 10.64898/2026.07.04.736150 medRxiv
Top 0.1%
8.1%
Show abstract

Parkinsons disease (PD) is a movement disorder characterized by the accumulation of alpha-synuclein aggregates leading to dopaminergic neuron loss in the substantia nigra. While PD has been associated with environmental and microbiome changes, our ability to assess the mechanistic impact of these factors on synuclein aggregation in cells has remained limited. Here, we designed and optimized a high-throughput optical screening system to assess the effect of metabolites and small molecules on synuclein aggregation in cell lines expressing a synuclein-fluorescent protein fusion and treated with pre-formed fibrils (PFFs). Using this assay, we identified several compounds that modulate synuclein aggregate accumulation in cells, including harman, a {beta}-carboline that led to reduced synuclein aggregation. We further investigated the transcriptional effect of harman and PFFs and identified changes in peroxiredoxins as a potential mechanism linking harman to aggregate accumulation. Altogether, this work establishes a pipeline to prioritize small molecules that can impact synuclein aggregate formation.

16
HTS-Oracle v2: Prospective AI-Guided Discovery and Experimental Validation of Small Molecule Modulators Across Multiple Targets

Abdel-Rahman, S.; Gabr, M.

2026-06-19 bioinformatics 10.64898/2026.06.15.732399 medRxiv
Top 0.1%
8.0%
Show abstract

High-throughput screening (HTS) remains the cornerstone of early-phase small molecule discovery yet consistently underperforms against immunotherapy targets, yielding validated hit rates below 0.1%. Here we introduce HTS-Oracle v2, which features rigorous cross-validation that ensures honest performance estimates. HTS-Oracle v2 was trained and validated across four clinically significant immune checkpoint targets (CD28, ICOS, LAG-3, and TIGIT) achieving ROC-AUC values of 0.968, 0.969, 0.875, 0.928 respectively under rigorous cross-validation. For prospective experimental validation, HTS-Oracle v2 was applied to an 8,960-compound Enamine Protein Mimetic Library, selecting only 25 compounds per target for experimental testing using temperature-related intensity change (TRIC) technology, a 99.7% reduction in screening burden. HTS-Oracle v2 identified 4, 5, 4, and 6 validated binders from 25 prospectively selected compounds per target, corresponding to validated hit rates of 16%, 20%, 16%, and 24%, respectively. Notably, 67-80% of all experimentally confirmed hits across the full 8,960-compound library were captured within just 25 model-selected compounds per target. For CD28, this represents a 28-fold improvement over HTS-Oracle v1 (239x versus 8.4x), establishing HTS-Oracle v2 as an efficient platform for AI-guided prospective hit discovery across immunotherapy targets.

17
THEOBROMA: an aggregated open database of 1.13 million natural products with per-compound license auditing, three-tier classification, and stereochemistry-aware deduplication

Klamt, T.; Jaczkowski, A.; Franke, J.; Nejdl, W.

2026-06-16 bioinformatics 10.64898/2026.06.12.731585 medRxiv
Top 0.1%
8.0%
Show abstract

Natural products remain one of the most productive sources of pharmacologically active compounds for drug discovery, yet the current open aggregator landscape attributes licenses at database rather than compound granularity, with consequences that have become tangible as the field grows. A recent relicensing event in one constituent source (the September 2024 transition of the Natural Products Atlas to CC BY-NC 4.0) demonstrates how database-level licensing propagates across an aggregate and motivates the per-compound audit framework presented here. The same peer cohort separately leaves classification provenance and stereoisomer-family relations coarser than either layer warrants. THEOBROMA, accessible at https://theobroma.l3s.uni-hannover.de, integrates 1,133,004 natural products from 29 open sources under a per-compound license audit that resolves each compounds license tier across all attesting sources under a most-restrictive-wins rule, identifying 900,170 compounds (79.4%) under open-use licenses and exposing the per-source attestation chain and resolved tier through a dedicated audit endpoint and a query-time license filter. A three-tier classification stratifies 89.3% coverage into 35.1% curated, 43.9% high-confidence inferred, and 10.3% exploratory tiers, with 486,215 stereoisomer families preserved by full 27-character InChIKey deduplication and exposed via a dedicated /api/stereoisomers/<comp_id> endpoint and a radial-family display. Per-compound license provenance is the primary differentiator. Classification stratification and stereoisomer-family exposure add finer-grained access to two related axes, supporting license-compatible virtual screening and isomer-specific bioactivity analysis at corpus scale. As an evolving open resource, THEOBROMA pairs continuous pipeline maintenance with interactive geographic, taxonomic, and chemical-space exploration.

18
A covalent irreversible inhibitor binds in two mutually exclusive conformations to the active-site cysteine residue of human aldehyde dehydrogenase 1A3

Covaleda, D.; Vizarraga, D.; Upadhyay, T.; Zhu, J.; Abegg, D.; Pequerul, R.; Hugo, M.; Adibekian, A.; Fita, I.; Pares, X.; Aviles, F. X.; Boggyo, M.; Farres, J.

2026-07-15 biochemistry 10.64898/2026.07.14.738401 medRxiv
Top 0.1%
7.9%
Show abstract

Aldehyde dehydrogenases (ALDH) are enzymes that catalyze the NAD(P)+-dependent oxidation of aldehydes into carboxylic acids, playing roles in detoxification, biosynthesis, and regulatory functions. Dysfunction of ALDH is associated with serious conditions such as alcohol intolerance, cancer, cardiovascular problems, and neurological disorders. In humans, ALDH1A1 and ALDH1A3 isoforms act as retinaldehyde dehydrogenases and are overexpressed in various cancers, where high levels are associated with increased tumor malignancy, cancer stem cell traits, and therapeutic resistance. ALDH1A3 is recognized as a promising target for anticancer therapies, with several inhibitors, mainly reversible, developed to specifically target it or the enzyme family. Since ALDH enzymes can also display esterase activity, we used this property to develop an in vitro assay specifically targeting the esterase function of ALDH1A3. A highly conserved active-site cysteine in ALDH1A3 is located at the bottom of two converging channels, which define the substrate- and cofactor-binding pockets. To target this catalytic cysteine, we screened a library of 3,200 cysteine-focused covalent fragments. This led to the identification of Z3405279217 (Z34), an acrylamide-based covalent compound that inhibits both ALDH1A1 and ALDH1A3 at sub-micromolar levels. Biochemical and biophysical tests confirmed that Z34 acts as a time-dependent, covalent, and irreversible binder to the active-site cysteine. In this work, we determined the Cryo-EM structure of the ALDH1A3-Z34 complex at 2.26 [A] resolution, confirming the covalent attachment to the catalytic cysteine of Z34. Notably, two mutually exclusive covalent binding modes were observed: one occupying the substrate-binding pocket and the other the cofactor-binding region. Z34 displayed unexpected binding modes within the active site and holds promise as a lead compound for future drug development. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=184 HEIGHT=200 SRC="FIGDIR/small/738401v1_ufig1.gif" ALT="Figure 1"> View larger version (33K): org.highwire.dtl.DTLVardef@982d1forg.highwire.dtl.DTLVardef@ba86f2org.highwire.dtl.DTLVardef@1f19f2borg.highwire.dtl.DTLVardef@8e807_HPS_FORMAT_FIGEXP M_FIG C_FIG

19
MAERM: Predicting Enzyme-Reaction Matching Relationships with a Mixed-Attention Model

Liu, T.; Zhai, S.; Lin, S.; Zhan, X.; Deng, J.; Liu, H.; Siu, S. W. I.

2026-07-10 bioinformatics 10.64898/2026.07.06.736902 medRxiv
Top 0.1%
7.8%
Show abstract

Harnessing enzyme specificity requires a thorough understanding of enzyme promiscuity, which determines enzymes catalytic scope; however, measuring this scope still relies heavily on labor-intensive analytical approaches. While data-driven approaches have emerged to predict the catalytic scope of enzymes, these methods continue to face challenges such as restricted datasets and insufficient integration of enzyme structural information and reaction transformations. Here, we introduce MAERM, an innovative mixed-attention model designed to predict enzyme-reaction matching relationships. Built on our MAERM-DB, a dataset with broad coverage of validated and chemoenzymatic catalysis data, MAERM utilizes a local-global attention module to integrate multimodal enzyme information with fine-grained reaction representations, thereby predicting enzyme-reaction matching probabilities. Results show that MAERM consistently outperforms all baselines, with an average F1-score of 0.984. Notably, on challenging test samples with less than 40% sequence identity to the training set, MAERM outperforms the second-ranked model by 5.9% in F1-score. In addition, MAERM achieves the highest top-10 success rate of 51.7% on Enzyme-405 and the highest balanced accuracy of 0.697 on BioCat-547, further supporting its generalizability in enzyme screening and chemoenzymatic catalysis. Finally, MAERM can serve as an efficient scoring module. When integrated with ProteinMPNN, MAERM has successfully guided novel enzyme design for two carbonyl reduction reactions, resulting in enhanced catalytic potential for the native substrate and demonstrating broad compatibility. Overall, MAERM has the potential to reduce the experimental cost of measuring enzymes catalytic scope, facilitate enzyme design, and ultimately accelerate the design-build-test-learn cycle in enzyme engineering.

20
MSAgent: An Evidence Grounded Agentic Framework for LLM-driven Scientific Exploration in Mass Spectrometry-based Metabolomics

Li, Y.; Zhong, Y.; Liu, P.; Yusheng, T.; Zhan, H.; Xia, J.

2026-04-24 bioinformatics 10.64898/2026.04.22.720103 medRxiv
Top 0.1%
7.7%
Show abstract

Mass spectrometry (MS) is a cornerstone high-throughput technology for molecular discovery, yet the reliable elucidation of chemical structures remains a formidable, expert-dependent bottleneck. Currently, achieving a reliable molecular identification from raw mass spectra necessitates a manual assembly--a labor-intensive ordeal of heuristic reasoning and the tedious integration of siloed computational tools, perpetuating a profound throughput gap between rapid data acquisition and the glacial pace of structural annotation. Here we present MSAgent, an autonomous agentic framework that bridges the gap between computational automation and expert intuition by emulating the cognitive logic of human specialists. By orchestrating a MSToolbox of over 50 domain-specific tools via Large Language Models (LLMs), MSAgent dynamically unifies the analytical pipeline into a scalable, evidence-grounded workflow, allowing for intent-aware planning, cross-resources outputs synthesis, and visual mechanistic interpretation within traceable reasoning chains and evidence-backed analytical reports. We evaluated MSAgent across multiple open benchmarks, including the established community challenges - Critical Assessment of Small Molecule Identification (CASMI) 2016/2022, CANOPUS, and LLM-oriented test cases. On CASMI, MSAgent consistently boosts retrieval performance by over 10% MRR across diverse benchmarks while ensuring high reliability--improving or preserving ranks in 95% of cases. For more challenging molecular de novo tasks on CANOPUS, MSAgent builds upon the outputs of baseline models with consistent refinement, yielding over a 40% average gain in Tanimoto similarity for ground-truth recovery. In addition, MSAgent demonstrates remarkable advantages in eliminating the hallucination phenomenon over LLMs without domain tool support, producing better-calibrated confidence (Pearson r = 0.438 vs -0.219 for gpt-4o). It improves exact-match rate by 38.8% over gpt-4o in candidate discrimination tasks, and achieved a 64% success rate in recommending high-quality candidate structures with Tanimoto similarity more than 0.7, where gpt-4o predominantly selected candidates with similarity below 0.3. Our work enables high-throughput mass spectrometry data to be analyzed in an intent-driven and automated manner, lowering the analysis barrier for no-expert to deliver molecular identification result with transparent analytical process, and accelerating discovery in metabolism and related fields by bridging the gap between experimental data acquisition and computational interpretation.